Papers by Kiet Van Nguyen

6 papers
ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization (2025.coling-demos)

Copied to clipboard

Challenge: ViSoLex is an open-source repository for Vietnamese lexical normalization . it provides two core services: Non-Standard Word (NSW) Lookup and Lexical Normalization enabling users to retrieve standard forms of informal language and standardize text containing NSWs.
Approach: They propose to integrate pre-trained language models and weakly supervised learning techniques to ensure accurate and efficient normalization.
Outcome: The system provides two core services: Non-Standard Word (NSW) Lookup and Lexical Normalization, enabling users to retrieve standard forms of informal language and standardize text containing NSWs.
ViHOS: Hate Speech Spans Detection for Vietnamese (2023.eacl-main)

Copied to clipboard

Challenge: Increasing use of social networking sites can cause problems for human moderators to review tagged comments.
Approach: They present a dataset that contains 26k spans on 11k comments and detailed annotation guidelines . they also provide definitions of hateful and offensive spans in Vietnamese comments .
Outcome: The proposed dataset shows that it is difficult to detect specific types of spans in the dataset . the dataset is the first human-annotated corpus containing 26k spans on 11k comments .
Revealing Weaknesses of Vietnamese Language Models Through Unanswerable Questions in Machine Reading Comprehension (2023.eacl-srw)

Copied to clipboard

Challenge: Existing problems in Vietnamese Machine Reading Comprehension systems are limited due to multilinguality, which limits the ability of multilingual models to develop state-of-the-art systems.
Approach: They propose to modify the process of annotating unanswerable questions to improve the quality of unanswered questions to a higher level of difficulty for Machine Reading Comprehension systems to solve.
Outcome: The proposed modification improves the quality of unanswerable questions to a higher level of difficulty for Machine Reading Comprehension systems to solve.
ViGoEmotions: A Benchmark Dataset For Fine-grained Emotion Detection on Vietnamese Texts (2026.eacl-long)

Copied to clipboard

Challenge: Recent advances in NLP have greatly improved outcomes in emotion prediction and harmful content detection.
Approach: They propose to classify Vietnamese comments into 27 distinct emotions using a model-based lexical normalization system and a transformer-based model.
Outcome: The proposed corpus of 20,664 social media comments is based on a novel model that can support multiple architectures, but its quality and preprocessing strategies remain key factors influencing performance.
A Large-Scale Benchmark for Vietnamese Sentence Paraphrases (2025.findings-naacl)

Copied to clipboard

Challenge: 1.2M original–paraphrase pairs were generated using a hybrid approach to generate high-quality paraphrases.
Approach: They present a high-quality Vietnamese dataset for sentence paraphrasing . they used automatic paraphrase generation and manual evaluation to ensure high quality .
Outcome: The proposed dataset is the first large-scale study on Vietnamese paraphrasing . it combines automatic paraphrase generation with manual evaluation to ensure high quality .
ViNLI: A Vietnamese Corpus for Studies on Open-Domain Natural Language Inference (2022.coling-1)

Copied to clipboard

Challenge: a large-scale corpus is needed for studies on natural language inference (NLI) for Vietnamese, which can be considered a low-resource language.
Approach: They propose a corpus for evaluating Vietnamese natural language inference models . they use a human-annotated corpus extracted from more than 800 online news articles .
Outcome: The ViNLI corpus is created and evaluated with a strict process of quality control . the best system performance is still far from human performance (a 14.20% gap in accuracy).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations